Papers with language generation models
Mitigating Societal Harms in Large Language Models (2023.emnlp-tutorial)
Copied to clipboard
| Challenge: | Recent studies have highlighted societal harms that can be caused by language generation models deployed in the wild. |
| Approach: | They propose to use a typology of technical approaches to mitigating harms of language generation models to provide an overview of potential social issues in language generation including toxicity, social biases, misinformation, factual inconsistency, and privacy violations. |
| Outcome: | The proposed typology addresses toxicity, biases, misinformation, factual inconsistency, and privacy violations in language generation models. |
PRAL: A Tailored Pre-Training Model for Task-Oriented Dialog Generation (2021.acl-short)
Copied to clipboard
| Challenge: | Existing approaches to building task-oriented dialog systems require a substantial amount of annotations and thus are labor-intensive. |
| Approach: | They propose a Pre-trainedRole Alternating Language model (PRAL) that is explicitly designed for task-oriented dialog tasks. |
| Outcome: | The proposed model outperforms or is on par with state-of-the-art models on task-oriented dialog tasks. |
MulZDG: Multilingual Code-Switching Framework for Zero-shot Dialogue Generation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing zero-shot dialogue generation systems rely on large-scale pre-trained language models. |
| Approach: | They propose a multilingual learning framework for zero-shot dialogue generation that can transfer knowledge from an English corpus to a non-English corpus with zero samples. |
| Outcome: | The proposed framework can transfer knowledge from an English corpus to a non-English corpus with zero samples. |
Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences (2021.emnlp-main)
Copied to clipboard
| Challenge: | aaron carroll: in social settings, human behavior is governed by unspoken rules of conduct rooted in societal norms . carroll and colleagues examine whether language generation models can serve as behavioral priors if they are not . they say we examine whether they can generate descriptions of actions that accomplish predefined goals . |
| Approach: | They propose to combine multiple expert models to improve quality of generated actions, consequences, and norms. |
| Outcome: | The proposed models significantly improve the quality of generated actions, consequences, and norms compared to baselines. |
Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation (2022.findings-acl)
Copied to clipboard
| Challenge: | Current language generation models suffer from issues such as repetition, incoherence, and hallucinations . |
| Approach: | They propose to analyze exposure bias from an imitation learning perspective and prove it is a problem . they show that exposure bias leads to an accumulation of errors during generation . |
| Outcome: | The proposed model fails to capture errors during generation and poor generation quality. |
Provably Confidential Language Modelling (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to train language models without memorizing sensitive data are mismatched and can be difficult to screen and filter. |
| Approach: | They propose a method to train language generation models while protecting the confidential segments of training data. |
| Outcome: | The proposed method prevents unintended memorization by randomizing parts of the training process while protecting strong confidentiality. |
On the Efficacy of Sampling Adapters (2023.acl-long)
Copied to clipboard
| Challenge: | Using sampling adapters can improve the quality of the generated text. |
| Approach: | They propose a framework for understanding sampling adapters and propose 'sampling adapters' they argue that the shift enforced by them can be viewed as a trade-off between precision and recall . |
| Outcome: | The proposed framework can be used to improve the quality of language models by modifying their distributions to improve their precision and recall. |
Robust Conversational Agents against Imperceptible Toxicity Triggers (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing work to generate adversarial attacks is costly and not scalable . despite the abundance of research in this area, little attention has been given to adversarials . |
| Approach: | They propose an adversarial attack mechanism that mitigates toxic language generation . they propose a defense mechanism that is scalable and can be generalized . |
| Outcome: | The proposed defense is effective at avoiding toxic language generation even against imperceptible toxicity triggers while preserving conversational flow. |
Language Generation Models Can Cause Harm: So What Can We Do About It? An Actionable Survey (2023.eacl-main)
Copied to clipboard
| Challenge: | Recent advances in the capacity of large language models to generate human-like text have prompted a heated discourse around the risks of societal harms they introduce. |
| Approach: | They propose a taxonomy of interventions organized around the different phases where they can be adopted to mitigate harms. |
| Outcome: | The proposed methods are based on several prior works’ taxonomies of language model risks and provide an overview of strategies for detecting and ameliorating different kinds of risks/harms. |
Adapting a Language Model for Controlled Affective Text Generation (2020.coling-main)
Copied to clipboard
| Challenge: | Existing models for affective text generation fail to capture emotional aspects of conversations without explicit affective information. |
| Approach: | They propose to incorporate emotion as prior for the probabilistic state-of-the-art text generation model such as GPT-2 and incorporate emotion into the model to ensure grammatical correctness. |
| Outcome: | The proposed model outperforms existing models in all intensities and is robust to human evaluations. |
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand (2022.naacl-main)
Copied to clipboard
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander Fabbri, Yejin Choi, Noah A. Smith
| Challenge: | Recent advances on models and metrics should benefit and inform each other, authors argue . bidimensional leaderboards allow for fast, accurate evaluation of language generation models . |
| Approach: | They propose a bidimensional leaderboard that tracks progress in language generation models and metrics for their evaluation. |
| Outcome: | The proposed leaderboards track progress in language generation models and metrics for their evaluation. |
Retrieval Enhanced Model for Commonsense Generation (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing frameworks for commonsense generation are lacking for pre-trained models. |
| Approach: | They propose a framework that uses concept matching to retrieve prototype sentences and trainable sentence retriever to enhance pre-training and fine-tuning. |
| Outcome: | The proposed framework achieves state-of-the-art on the large-scale Common-Gen benchmark. |
Detecting Bot-Generated Text by Characterizing Linguistic Accommodation in Human-Bot Interactions (2021.findings-acl)
Copied to clipboard
| Challenge: | Language generation models' democratization makes it easier to generate human-like text at-scale for nefarious activities, from spreading misinformation to targeting specific groups with hate speech. |
| Approach: | They propose to use linguistic alignment to detect bot-generated text rather than using it directly. |
| Outcome: | The proposed methods are more robust across datasets and models if they use information about how people respond to it rather than using the bot's text directly. |
Twist Decoding: Diverse Generators Guide Each Other (2022.emnlp-main)
Copied to clipboard
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, Noah A. Smith
| Challenge: | Using a variety of language generation models, ensembling models is challenging during inference. |
| Approach: | They propose a method that decodes text models that do not assume a shared vocabulary, tokenization or generation order. |
| Outcome: | The proposed method outperforms models decoded in isolation over various scenarios. |
Evaluation of African American Language Bias in Natural Language Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that large language generation models disadvantaging African American Language (AAL) can be biased for certain language varieties, but there is little research on the impact of these biases on other languages. |
| Approach: | They evaluate how well LLMs understand African American Language (AAL) in comparison to white Mainstream English (WME) using a dataset of AAL texts from a variety of regions and contexts, they find dialectal bias in six pre-trained LLM. |
| Outcome: | The proposed models understand African American language in comparison to white mainstream English (WME) the proposed models have performance gaps on two tasks that are not matched by the model. |
NEUROSTRUCTURAL DECODING: Neural Text Generation with Structural Constraints (2023.acl-long)
Copied to clipboard
| Challenge: | Current approaches for conditional text generation focus on lexical constraints, but lack syntactic constraints to support complex semantic constraints. |
| Approach: | They propose a decoding algorithm that incorporates syntactic constraints to improve the quality of the generated text. |
| Outcome: | The proposed method improves on three different language generation tasks and shows improved lexical and syntactic metrics. |
Generalized Entropy Regularization or: There’s Nothing Special about Label Smoothing (2020.acl-main)
Copied to clipboard
| Challenge: | Prior work has explored regularizing the output distributions of probabilistic models to alleviate overfitting. |
| Approach: | They propose a family of entropy regularizers that have a connection to regularization . they find that label smoothing provably does not allow for sparsity in an output distribution . |
| Outcome: | The proposed method improves the relationship between model entropy and performance on language generation tasks. |
On Improving Summarization Factual Consistency from Natural Language Feedback (2023.acl-long)
Copied to clipboard
| Challenge: | Recent work shows that language generation models can make errors on fine-grained qualities such as factual consistency. |
| Approach: | They propose to use natural language feedback to improve generation quality and user preference alignment. |
| Outcome: | The proposed model can provide factual consistency in human-edited summaries and further insights into summarization factual consistentness. |
MASIVE: Open-Ended Affective State Identification in English and Spanish (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models that fail to understand cultural and language influences the meaning of emotional terms like "love" a new study shows that smaller finetuned models outperform much larger LLMs on region-specific span prediction tasks. |
| Approach: | They propose to use a reddit reddits dataset to identify a set of affective states . they find that smaller finetuned multilingual models outperform larger LLMs . |
| Outcome: | The proposed model outperforms larger models on span prediction task even on region-specific Spanish affective states. |